Back

Genomics, Proteomics & Bioinformatics

Preprints posted in the last 30 days, ranked by how well they match Genomics, Proteomics & Bioinformatics's content profile, based on 188 papers previously published here. The average preprint has a 0.11% match score for this journal, so anything above that is already an above-average fit.

1
TBpop: an open-access genomic portal integrating genomic variation, population genetic statistics, phylogeny, pangenome composition, and strain metadata of epidemic Mycobacterium tuberculosis strains from China

Zhou, Y.; Huang, F.; Zhao, Y.

2026-08-12 bioinformatics 10.64898/2026.08.06.742672 medRxiv
Top 0.1%
13.1%
Show abstract

Tuberculosis remains a major global public health threat. While whole-genome sequencing has transformed our understanding of the causative agent, Mycobacterium tuberculosis (MTB), existing genomic databases are highly fragmented and often underrepresent structural variations (SVs). Furthermore, critical population-genetic statistics are rarely integrated with phylogenetic and geographic context, forcing researchers to reconcile separate datasets manually. To address this gap, we developed TBpop (https://tbpop.chinacdc.cn), an open-access, integrated population genomics portal. TBpop is built from 420 clinical MTB isolates selected from the first national drug resistance baseline survey in China. The portal integrates isolate metadata, pangenome categories, SNPs, SVs, IS6110 insertion sites, strain phylogeny, and gene-level statistics, and provides three interactive explorer modules: the Population Explorer, the Statistics Explorer, and the Variation Explorer. Additionally, a User Analysis module allows researchers to run population genetic workflows on their own alignments. TBpop provides an integrated platform for exploring genome plasticity, signatures of positive selection, and conservation patterns of functionally important genes in MTB.

2
Predicting Early MASLD-HCC from Serum N-Glycomics: A SHAP-Interpreted Gaussian Naive Bayes Model Built on nLC-HCD-PRM-MS/MS Profiling

Lin, Y.; Chithravel, V.; Dai, J.; Liu, S.; Lubman, N. Y.; Lubman, D. M.

2026-08-06 gastroenterology 10.64898/2026.08.04.26359485 medRxiv
Top 0.1%
12.2%
Show abstract

Hepatocellular Carcinoma (HCC) arising from Metabolic Dysfunction-Associated Steatotic Liver Disease (MASLD) is an increasing public health burden with high mortality, highlighting the need for improved early detection strategies. Current surveillance tools, including Alpha-fetoprotein (AFP) and ultrasound, lack sufficient sensitivity for early-stage HCC detection. We analyzed serum samples from 131 patients, including 58 with cirrhosis and 73 with MASLD-related HCC (42 early-stage, 31 late-stage), using an nLC-stepped HCD-PRM-MS/MS workflow for targeted N-glycome profiling of glycopeptides derived from haptoglobin and vitronectin. Combining targeted glycopeptides with AFP significantly improved HCC detection compared with AFP alone. The optimal panel for all HCC versus cirrhosis (AFP + VTNC_169_A2G2F0S1 + VTNC_242_A3G3F2S2) achieved an AUC of 0.859 and 76.7% sensitivity at 90% specificity. For early-stage HCC, AFP + HP_184_A3G3F1S3 + VTNC_169_A2G2F0S1 yielded an AUC of 0.890 with 66.7% sensitivity at 1% specificity. A SHAP-selected Gaussian Naive Bayes model based on seven molecular/glycopeptide features, without demographic variables, further improved performance, achieving ROC-AUC values of 0.9985 in training and 1.0000 in independent testing cohorts, with accuracies of 98.1% and 100.0%, respectively.

3
Exploring vulnerable proteins in the progression of head and neck squamous cell carcinoma

Agrawal, A.; Kumar, S.; Vindal, V.

2026-08-13 bioinformatics 10.64898/2026.08.07.743269 medRxiv
Top 0.2%
9.2%
Show abstract

A protein whose removal or deletion causes significant disruption or collapse of a protein-protein interaction (PPI) network is referred to as a vulnerable protein. Such proteins may serve as valuable therapeutic or diagnostic targets in disease-associated networks. In this study, two PPI networks were constructed, one for HPV-positive and the other for HPV-negative head and neck squamous cell carcinoma (HNSCC), and the vulnerable proteins of these networks were identified by the node deletion approach. After analyzing the networks, 27 unique vulnerable proteins in HPV-positive and 72 unique vulnerable proteins in HPV-negative HNSCC were identified. Among them, one HPV-positive and seven HPV-negative HNSCC vulnerable proteins were further chosen by integrating multi-omics data. To exploit the vulnerabilities of these proteins, candidate synthetic lethal (SL) partners were predicted whose inhibition may selectively impair tumor survival. Subsequently, drug-gene interaction analysis was performed to identify inhibitors targeting the SL partners of these vulnerable proteins. Notably, in HPV-positive HNSCC, TOP2A, CHEK1, and CHEK2 genes were identified as SL partners of TTN, and their inhibitors were already clinically approved. While in HPV-negative HNSCC, ADA and MMP19 were identified as an SL partner of LMO7; TMEM45B, CDH3, and ELF3 genes were identified as an SL partner of CGN; and ZNF433 was identified as an SL partner of FLNC. However, MMP19, ZNF433, and TMEM45B inhibitors were not reported. Thus, these vulnerable proteins, including their SL partners, provide novel avenues to explore and develop more efficient and precise therapeutic and diagnostic strategies.

4
A meta-analysis of ancient and present-day Central Eurasian genome data to revise archaic hominin ancestry

Rymbekova, A.; Kuhlwilm, M.

2026-08-14 genomics 10.64898/2026.08.10.743976 medRxiv
Top 0.2%
7.3%
Show abstract

Archaic introgression has shaped the evolutionary history of Eurasian populations, yet Central Eurasian region remains understudied despite being at the crossroads of ancient human migration. Here, we analyzed the whole-genome data of five Central Eurasian (CE) individuals from Early Bronze Age (EBA) and five present-day CE individuals to characterize the archaic introgression landscape. We estimated that archaic introgression from Neanderthal and Denisovan archaic hominins comprises approximately 2.2% of the Central Eurasian genomes. Both amount and chromosomal distribution of archaic introgression remained largely unchanged between the EBA and present-day CE individuals. Putative introgressed fragments matching the Altai Neanderthal and the Altai Denisovan were retrieved. Our results suggest that while the archaic introgression levels seemingly remained stable over the past several thousand years, larger modern CE genomes panels will be required to fully characterize the genomic landscape of archaic ancestry in the region.

5
Multimodal spatial-omics reveal the heterogeneity and intercellular network characteristics of papillary craniopharyngiomas.

Jiang, Y.; Luo, H.; Zheng, H.; Li, C.; Zan, X.; Xu, J.; Chen, Y.

2026-08-24 cancer biology 10.64898/2026.08.20.746031 medRxiv
Top 0.3%
6.0%
Show abstract

Despite significant advancements in microsurgical techniques in recent years, the treatment and prognosis of craniopharyngiomas remain unsatisfactory. As a central nervous system tumor located adjacent to important brain structures such as the hypothalamus-pituitary axis and accompanied by a highly inflammatory microenvironment, the tumor heterogeneity and tumor microenvironment characteristics of papillary craniopharyngiomas (PCPs) remain unclear. In this study, we integrated multimodal single-cell and spatial profiling from PCP tissue and peripheral blood mononuclear cells (PBMCs) to elucidate the tumor heterogeneity and microenvironment characteristics of PCP. Our single-cell and spatial analyses defined four specific tumor cell states in PCP, representing specific transcriptional regulatory programs and spatial heterogeneity characteristics during tumor progression. By constructing a spatial niche composed of tumor, immune, and stromal cells, we analyzed the cellular and spatial ecosystem of PCP at multiple levels to further assess the communication relationships between different tumor cell states and microenvironment cells. This study established a multidimensional molecular atlas of PCP from the perspectives of cell state, spatial structure, and microenvironment interactions, providing a foundation for understanding its biological behavior and exploring new intervention strategies.

6
The Yamanashi Multi-omics Cohort (YMoC): study design of a screening-defined longitudinal metabolic-risk cohort with integrated multi-omics and digital phenotyping

Goto, G.; Hanawa, D.; Naito, K.; Wang, Q. S.; Kanai, S.; Awaji, M.; Nishikawa, H.; Yui, H.; Nishitani, S.; Miyake, K.; Ooka, T.

2026-08-21 epidemiology 10.64898/2026.08.18.26360529 medRxiv
Top 0.3%
5.6%
Show abstract

Background: Large-scale biobanks have advanced genomic and epidemiologic research, but many rely on infrequent biological sampling and limited digital phenotyping. The Yamanashi Multi-omics Cohort (YMoC) was established to support longitudinal assessment of molecular, clinical, and behavioural changes in a screening-defined cohort of adults at elevated metabolic risk without diagnosed diabetes. Methods: YMoC is a longitudinal cohort of 215 adults aged 30-70 years in Yamanashi Prefecture, Japan, who met prespecified glycaemic eligibility criteria at health check-up, including fasting plasma glucose 100-125 mg/dL (5.6-6.9 mmol/L) and HbA1c <6.5%. Participants underwent three in-person visits over six months. Measurements include 75-g oral glucose tolerance testing with serial sampling, clinical biochemistry, anthropometry, liver elastography, and collection of blood, urine, stool, and saliva for multi-omics profiling. Between visits, participants wore a Fitbit Inspire 3 and completed daily app-based questionnaires using the Taohealth app. Current molecular data include genome-wide single nucleotide polymorphism array genotyping and longitudinal plasma proteomics in a subset. Conclusions: YMoC is designed to evaluate within-person molecular and phenotypic trajectories in a screening-defined metabolic-risk cohort. The cohort provides a dense longitudinal resource linking clinical assessments, biospecimens, omics assays, and digital phenotyping, including analyses of insulin-resistance-related markers such as homeostasis model assessment of insulin resistance (HOMA-IR).

7
Relational Graph Convolutional Networks for Glioblastoma Biomarker Discovery via ceRNA and Copy Number Variation Analysis

Khandelwal, S.; Jarvis, N.; Zhan, J.

2026-08-20 bioinformatics 10.64898/2026.08.16.744525 medRxiv
Top 0.4%
4.7%
Show abstract

Glioblastoma (GBM) is a highly aggressive brain tumor with an extremely poor 5-year survival rate of 6.9%, largely attributable to the lack of reliable biomarkers. While competing endogenous RNA (ceRNA) and copy number variation (CNV) analyses offer unique biomarker identification potential, current approaches neglect the integration of multiple regulatory mechanisms for biomarker detection. To address this limitation, we applied relational graph convolutional networks (RGCNs) to ceRNA and CNV knowledge graphs through a novel late fusion ensemble architecture. The proposed architecture outperformed baseline models and identified five novel biomarkers, including hsa-miR-196a and hsa-miR-224. Kaplan-Meier survival analysis and Cox regression indicated that the identified genes hold significant prognostic and diagnostic power. The early stratification of the Kaplan-Meier curves indicates the potential these genes hold for patient survival prediction. The results illustrate that a late fusion RGCN ensemble effectively captures complex gene interactions, overcoming limitations of existing models and providing a framework for biomarker discovery. The novel biomarkers serve as prospective targets for future GBM therapeutic development and candidates for non-invasive diagnostic assays.

8
Cross-attention and language models reveal the interpretability of functional predictions for the human olfactory receptor family

Zhang, Y.-F.; Xu, Z.-h.; Gao, C.-x.; Duan, S.-Y.; Li, G.; Xu, C.; Lu, H.-M.

2026-08-18 bioinformatics 10.64898/2026.08.10.744067 medRxiv
Top 0.4%
4.5%
Show abstract

The attention mechanism offers the possibility for data-driven discovery of biological principles. However, for important protein families such as human olfactory receptors, the extent to which attention can associate with biologically meaningful key regions lacks systematic validation. In this study, using human olfactory receptors (ORs) as a model, we constructed CrossVOI, a VOC-OR interaction prediction framework based on protein language models and cross-attention, achieving predictive performance superior to existing methods. Furthermore, we systematically analyzed the attention distributions of CrossVOI and found that attention not only focused on ligand-binding interfaces and evolutionarily conserved sites, but also to some extent identified certain dynamically regulated regions. In summary, we propose CrossVOI, currently the best-performing framework for VOC-OR interaction prediction, and analyze the interpretability of the attention mechanism for human ORs. This study provides insights into the interpretability of protein function prediction methods and is expected to contribute to the exploration of attention mechanisms in biological mechanisms, and provide assistance for large-scale screening and mechanistic analysis of olfactory receptors.

9
RNA m6A demethylase ALKBH4 governs whitefly development

Yang, J.; Wang, C.; Gong, P.; He, C.; Fu, B.; Liu, S.; Wei, X.; Yin, C.; Huang, M.; Du, T.; Liang, J.; Zhou, X.; Nauen, R.; Zhang, Y.; Bass, C.; Yang, X.

2026-08-10 zoology 10.64898/2026.08.08.743705 medRxiv
Top 0.5%
4.2%
Show abstract

N6-methyladenosine (m6A) modification is the most predominant and ubiquitous internal modification of RNA in eukaryotes, serving as a key post-transcriptional regulator of gene expression that is dynamically modulated by methyltransferases (writers) and demethylases (erasers). However, while the functions of m6A methylases have been partially elucidated in insects, the identity of m6A erasers in arthropods and their chemical catalytic mechanisms, as well as biological functions, remains largely enigmatic. Here, we uncovered 2499 putative methylase genes and 1148 putative demethylase genes in 266 insect genomes, and demonstrated that ALKBH4 functions as an m6A demethylase in the whitefly, Bemisia tabaci, catalyzing the oxidative reversal of mRNA m6A modifications both in vitro and in vivo. Furthermore, we established that ALKBH4, in coordination with other core components of the m6A pathway, fulfills an essential function in regulating the transcript stability of Imaginal Disk Growth Factor 1 (IDGF1) during whitefly development. Collectively, our findings expand the evolutionary scope of the eukaryotic m6A modification system, and reveal a conserved yet insect-specific epitranscriptomic regulatory mechanism governing fundamental physiological processes and adaptive phenotypes. Significance statementThe addition of a methyl group to the N6-position of adenosine (m6A) is a highly abundant chemical modification of RNA. However, the functional role of m6A in insects and the key enzymes that regulate its levels remains poorly understood. In this study, we explored putative methylase genes and demethylase genes in hundreds insect genomes, and identified an m6A RNA demethylase, ALKBH4, in the whitefly, Bemisia tabaci. We demonstrate that ALKBH4 oxidatively reverses mRNA methylation in vivo and in vitro, in combination with other components of the m6A pathway, plays an important role in whitefly development. These findings provide new insight into m6A methylation system of insect.

10
Precision Transfusion Management: Rh Phenotype Compatibility and Antibody Surveillance in Southern China

Huang, X.-q.; Li, L.-x.; Yang, Z.-Y.; Long, X.-X.; Lai, C.-Y.

2026-08-10 hematology 10.64898/2026.08.05.26359788 medRxiv
Top 0.5%
4.1%
Show abstract

Objective: To investigate the distribution frequencies of Rh blood group antigens (C, c, D, E, e) and phenotypes in the population of Hengyang, Hunan Province, and to analyze the production of Rh alloantibodies in repeatedly transfused patients, thereby providing a basis for developing precise transfusion strategies. Methods: Rh phenotyping, antibody screening, and antibody identification were performed on 3,635 hospitalized patients and 5,326 blood donors using Rh blood group typing cards. A blood transfusion management system was used to identify and track patients' historical specific antibodies, with automatic alerts for inconsistent results. Results: The antigen frequency distribution in patients was D (99.56%) > e (94.69%) > C (91.64%) > c (48.06%) > E (38.79%). The phenotypic distribution frequencies among Rh(D)-positive patients were as follows: CCDee (51.31%) > CcDEe (30.01%) > CcDee (9.37%) > ccDEE (5.00%) > ccDEe (2.79%) > CCDEe (0.80%) > ccDee (0.39%) > CcDEE (0.28%) > CCDEE (0.05%). From March to October 2023, after implementing Rh phenotyping and antigen-matched compatible transfusions for five antigens, the antibody screening positivity rate decreased to 0.97%, compared to 1.14% during the same period in 2022 (p < 0.05). Antibody identification in 276 antibody-positive samples revealed that alloantibodies against the Rh system accounted for the highest proportion (46.01%, 127/276), which was lower than the 55.21% observed in 2022 (p < 0.05). Unexpected antibodies in the Rh system were the primary cause of crossmatch incompatibility in clinical transfusions, accounting for 46.01%. Conclusion: Rh phenotyping and sustained antigen-matched compatible transfusions in repeatedly transfused patients can effectively prevent and reduce alloantibody production. Continuous tracking of specific antibodies and transfusion efficacy evaluation can be achieved through an efficient blood transfusion management system.

11
A Practical Framework for Constructing Population-Specific and Alternate-Contig-Aware Genome References: A case study of Vietnam

Vo, N. S.; Tran, T. T. H.; Duong, V. C.; Nguyen, N. N.; Pham, T. M.; Vu, Q. T.; Tran, M. H.; Hoang, T. H.; Nguyen, Q.; Nguyen, D. T.

2026-08-27 genomics 10.64898/2026.08.24.746817 medRxiv
Top 0.5%
4.1%
Show abstract

Current studies in human genomics typically rely on the standard genome reference GRCh38 which is known to be biased toward populations of European ancestry and therefore has limitations when applied to other populations. Although various graph-based pangenome references were constructed for several populations to deal with this bias, their usage in practice is currently still limited compared to linear genome references. Here we present a framework for constructing a population-specific genome reference using GRCh38 as backbone with alternate-contig awareness to enhance genomic data analysis in the target population. We demonstrated the advantages of our framework using both public and in-house Vietnamese whole-genome sequencing (WGS) datasets. Genomic variants derived from high-coverage WGS data of the 1000 Vietnamese Genomes Project (VN1K) were imported into our framework to build a Vietnamese-specific Genome Reference (VGR). VGR was then compared to GRCh38 in read alignment and variant calling using high-coverage WGS data of 99 Vietnamese individuals (KHV) from the 1000 Genomes Project (1kGP). Using Omni array genotyping data from 99 KHV samples as an independent benchmark, we found that VGR improved variant-calling precision and reduced false-positive calls compared to GRCh38. Our framework could be easily used for other populations as long as they have a variant database similar to VN1K. Our code is publicly available at github.com/VinGenome/VGR

12
Cross-Species Comparison of Topologically Associating Domains (TADs) in Cereals Reveals Their Role in Genome Stability During Evolution

Li, E.; Huang, L.; Shi, J.; Xu, G.; Liu, H.; Jin, W.; Wang, Y.; Tang, S.; Diao, X.; Song, W.; Xin, B.; Lai, J.; Chen, J.

2026-08-19 plant biology 10.64898/2026.08.13.744466 medRxiv
Top 0.5%
4.0%
Show abstract

Topologically associating domains (TADs) are essential structural and functional modules of the genome that play a crucial role in regulating gene expression. In this study, we systematically investigated the conservation and evolution of TADs in five closely related crops, including maize, sorghum, coix, foxtail millet and broomcorn millet. Our results show that 74% of TAD boundaries are conserved between two inbred maize lines, B73 and Mo17, and that approximately 50% or more of TAD boundaries are conserved across different crop species. TAD number remains relatively stable in the face of changes in genome size. However, the length of TADs varies depending on genome size. Furthermore, we found that large-scale transposable element expansion leads to TAD expansion, while chromosomal inversions lead to TAD fusion and the formation of new TAD boundaries. Frequent chromatin interactions between subgenome chromosomes occur after whole-genome duplication. Moreover, we also found that crossovers are enriched at TAD boundaries in maize, indicating the importance of TADs as a fundamental unit during species evolution. Overall, our study provides insights into the conservation and evolution of TADs in crop genomes and their roles in genome organization and function.

13
Prediction of plant organismal complexity based on transcription factor annotation: an AI approach

Varshney, D.; Tajjar, M. H.; de Vries, J.; Hutter, F.; Rensing, S. A.

2026-08-22 evolutionary biology 10.64898/2026.08.18.745462 medRxiv
Top 0.6%
3.2%
Show abstract

How morphological complexity evolves is still enigmatic. While there is evidence in algae and plants as well as animals that diversification of the repertoire of transcription factors (TF) is causative for evolution of organismal complexity, there are many examples from lineages that follow their own way of complexity evolution, for example by expansion of particular families. For land plants, correlation of the size of the TF complement with number of cell types (as a proxy for morphological complexity) has been shown, and several families were identified as candidates to drive complexity evolution. Here, we expand a previously available dataset of cell type numbers from 12 to 82 proteomes and introduce a four class body plan scheme. We find that the total TF complement correlates with the number of cell types of Archaeplastida (primary plastid bearing plants and algae). We used TabPFN (Tabular Prior-data Fitted Network) for binary (uni- vs. multicellularity) as well as for four class Bauplan classification. TabPFN is able to predict the morphological complexity with high accuracy. This approach allows to determine organismal complexity based on the gene space of an organism. Based on our results, we can confirm that plant morphological evolution is driven by gain and expansion of TF families.

14
Transcriptomic Analysis Identifies Transient Mesendodermal State and Lineage Divergence in Human Pluripotent Stem Cell Differentiation

Borges, A. C.; Branco, M. A.; Cotovio, J. P.; Gomes, A. R.; Saraiva, J. E.; Moreira, L. M.; Cabral, J. M. S.; Henrique, D.; Diogo, M. M.; Fernandes, T. G.

2026-08-25 bioengineering 10.64898/2026.08.24.746642 medRxiv
Top 0.7%
3.1%
Show abstract

Human pluripotent stem cells serve as a vital model for studying early human lineage specification, yet conventional assessments relying on endpoint canonical markers of the three germ layers may overlook transient intermediate states and broader cellular programs. Here we combined directed differentiation of human induced pluripotent stem cells toward neuroectodermal, cardiac mesodermal, and hepatic endodermal lineages with comparative transcriptomic profiling across timepoints. Our analyses revealed a transient primitive streak-like mesendodermal state shared by mesodermal and endodermal trajectories, followed by lineage-specific divergence characterized by distinct transcriptional, metabolic, proliferative, and chromatin remodeling dynamics. Notably, endodermal differentiation exhibited rapid definitive endoderm commitment with enriched oxidative metabolism, whereas cardiac mesoderm differentiation showed progressive transcriptional remodeling and cardiac progenitor activation. These findings demonstrate that comparative transcriptomics can resolve developmental intermediates and cellular-state dynamics during human germ layer specification, providing a framework for evaluating lineage commitment beyond endpoint canonical marker expression, and to inform strategies for optimizing or redirecting differentiation.

15
Ori-Finder-Arch: An Updated Web Server for the Annotation and Visualization of Archaeal Replication Origins

You, Z.; Zhang, Z.; Luo, H.; Gao, F.

2026-08-19 bioinformatics 10.64898/2026.08.15.744077 medRxiv
Top 0.7%
3.1%
Show abstract

Archaea are promising chassis organisms in biotechnology, and the accurate annotation of their chromosomal replication origins (oriCs) is the key to unlocking their full potential. However, the existing Ori-Finder 2 web server suffers from low accuracy, slow speed, and limited scalability. In this study, we present Ori-Finder-Arch, an updated web server for high-performance oriC prediction in archaea. This pipeline integrates HMMER-based replication initiation protein (RIP) annotation, refined consensus motif recognition, and GC profile-based DNA unwinding element (DUE) detection. On a benchmark set of experimentally validated oriCs, Ori-Finder-Arch achieved a recall of 95.6% and a precision of 86.0%, substantially outperforming Ori-Finder 2 (62.2% and 63.6%, respectively), while running 4.75 times faster and supporting diverse assembly levels. When applied to the available archaeal assemblies, it successfully annotated 17,472 oriCs. Meanwhile, the web server provides interactive visualizations at different levels. In conclusion, Ori-Finder-Arch offers an efficient, accurate, and user-friendly platform for advanced studies of archaeal DNA replication initiation and synthetic biology applications, and is freely available at https://tubic.org/Ori-Finder-Arch/ and https://tubic.tju.edu.cn/Ori-Finder-Arch/.

16
Bridging Morphology and Genomics: A rapid image-based assessment of genomic admixture in the endangered gayal (Bos frontalis)

Ma, J.; Chen, Y.; Guo, Z.; Xiao, J.; Wu, H.; Luo, J.; Zhang, Y.-p.; Li, Y.

2026-08-25 zoology 10.64898/2026.08.25.746947 medRxiv
Top 0.7%
3.0%
Show abstract

Abstract The gayal (Bos frontalis) is an endangered semi-domesticated bovine species renowned for its high-quality beef. However, its semi-feral lifestyle, ongoing habitat fragmentation, and extensive genetic introgression from sympatric local cattle have led to dramatic population decline and severe erosion of purebred genetic integrity, posing substantial challenges to its conservation and utilization. To address the urgent demand for rapid, non-invasive, and field-compatible germplasm identification, we developed an integrated artificial intelligence (AI) framework that predicts genomic admixture composition from external morphological images. We constructed a comprehensive dataset comprising 6,245 morphological images and matched genomic sequences from 52 gayals maintained at the Yunnan Provincial Gayal Conservation Farms. Following a preliminary evaluation of nine deep learning models, five were incorporated into a anatomical segment-based multi-modal pipeline, among which Inception_V3 delivered the optimal overall performance. To enhance simultaneous extraction of local fine-grained features and global structural information, we further designed an innovative HybridInceptionViT model by integrating the multi-scale Inception module with the Vision Transformer (ViT) framework. This hybrid model significantly outperformed the baseline Inception_V3, boosting the accuracy of phenotype-derived prediction against genomic admixture estimate from 69.69% to 87.87% (absolute error <15%). This study establishes a practical, low-cost "phenotype-to-genotype" tool for rapid on-site gayal germplasm screening, offering a scalable strategy for the conservation and breeding management of endangered livestock, and holds broad application prospects for agricultural and livestock production systems.

17
Causally-inspired meta-representation learning framework for predicting patient-specific clinical responses to drug combinations

Zhang, Q.-Q.; Zhang, S.-W.; Shi, M.-H.; Li, J.-N.; Qiang, Y.-R.; Zhang, T.-H.

2026-08-21 bioinformatics 10.64898/2026.08.13.744613 medRxiv
Top 0.8%
2.6%
Show abstract

Large-scale prediction and assessment of clinical patient responses (i.e., RECIST class) to drug combinations remains challenging due to scarce patient-derived data. The existing prediction methods mainly rely on cancer cell line models. However, substantial biological heterogeneity between cancer cell lines and cancer patients within same tissues, as well as the heterogeneity between one tissue and another, often limit the generalizability of these methods in clinical patients. To overcome these limitations, here we present CaMeRe, a Causally-inspired Meta-representation learning framework designed to predict patient-specific clinical Response to drug combinations. In situations where stable causal factors and domain-specific response-modulating factors are unobservable, explicit discrete domain labels are unavailable, and data is scarce, CaMeRe designed a domain-invariant causal representation learning (DICRL) model guided by the invariant information bottleneck theory and causal intervention invariance principle, and also built a meta-learning framework with bi-level domain generalization to optimize DICRL model for achieving multi-domain generalization within and across tissues. By integrating the causal representation learning and meta learning framework, CaMeRe not only exhibited robust multi-domain generalization performance across multiple clinical drug combination response datasets and PDXs drug combination response datasets and generalization scenarios, but also had better interpretability. We applied CaMeRe to predict drug-combination response scores for 3,423 patients across 542,080 drug combinations. The predicted scores were significantly associated with biomarkers of known drug combinations and enabled the prioritization of candidate drug combinations across 11 cancer types, with stronger support from literature and clinical trial evidences than random baselines. We believe that CaMeRe can be a useful tool for predicting large-scale clinical individual drug combination responses and it has broad clinical applications.

18
SALRR: Scalable Analysis of Long-Read RNA-Seq Enables Comprehensive Transcriptome Profiling in Human Brain

Kouam, C.; Mingle, J.; Alvarez Jerez, P.; Evans, A.; Moller, A.; Baker, B.; Weller, C.; Paquette, K.; Brooks, J.; Grant, S. M.; Ayuketah, A.; Meredith, M.; Palade, J.; Malik, L.; Hise, K.; Raphael Gibbs, J.; Anderson, J.; Ding, J.; Harbert, R.; Fu, Y.; Zheng, X.; Garcia-Ruiz, S.; Gustavsson, E. K.; Blauwendraat, C.; Ryten, M.; Sedlazeck, F.; Ferrucci, L.; Reed, X.; Nalls, M. A.; Cookson, M. R.; Van Keuren-Jensen, K.; Hutchins, E.; Jain, M.; Billingsley, K. J.

2026-08-29 genomics 10.64898/2026.08.27.747499 medRxiv
Top 0.8%
2.6%
Show abstract

Isoform-resolved transcriptomics is fundamental to decoding the molecular complexity of the human brain, yet population-scale long-read RNA sequencing has remained inaccessible due to labor-intensive library preparation, sensitivity to RNA degradation in postmortem tissue, and the absence of integrated, reproducible analysis pipelines. Here we present SALRR (Scalable Analysis of Long-Read RNA-seq), an integrated wet-lab and computational platform designed to overcome these barriers. Automated ONT long-read cDNA library preparation on the Hamilton Microlab NGS STAR platform reduces hands-on time by 67% and enables 24 libraries per operator per day while maintaining performance across RNA integrity values. A modular, Snakemake-based pipeline performs end-to-end processing from ONT signal data to isoform-level quantification, incorporating SIRV spike-in calibration, multi-stage quality control, and stringent isoform validation. Applied to 10 postmortem frontal cortex samples from the North American Brain Expression Consortium, SALRR identified 31,607 high-confidence isoforms from 10,075 genes, including 8,532 novel splice variants absent from GENCODE v49, and complex splicing events systematically missed by short-read sequencing at neurodegeneration-relevant loci, including GBA1, CCNF, CHCHD10, and TREM2. All protocols and code are openly available, providing a scalable, community-ready framework for isoform-resolved transcriptomics in neurodegeneration, aging, and complex brain disease.

19
Predicting Protein-RNA Binding Affinity Changes via Spatial Coupling-Aware State Space Modeling

Chen, R.; Huang, X.; Jiang, H.; Ma, W.; Bi, X.; Wei, Z.; Nie, J.; Zhang, S.

2026-08-24 bioinformatics 10.64898/2026.08.23.745486 medRxiv
Top 0.8%
2.5%
Show abstract

Accurately predicting the effects of mutations on protein-RNA binding is crucial for elucidating disease mechanisms. Yet, exhaustively exploring the space of all possible variants is prohibitively expensive, motivating computational methods that can quantify mutation-induced changes in binding affinity (aka {Delta}{Delta}G) accurately and efficiently. We present iSCALE, an interpretable and generalizable deep learning method that adopts an implicit Spatial Coupling-Aware Ligand Encoding strategy to predict mutation-induced binding affinity changes. By injecting this implicit multiscale encoding scheme into a bidirectional state space modeling architecture, iSCALE learns a generalizable multiscale coupling pattern that achieves superior performances on not only the protein-RNA binding {Delta}{Delta}G, but also the protein stability {Delta}{Delta}G and protein-protein binding {Delta}{Delta}G predictions. Detailed analyses demonstrate that the model attention scores align well with structural characteristics. In addition, iSCALE shows good discriminative ability when predicting close samples such as complexes of same mutation but with different ligands or the same complex but with different mutation sites. In summary, iSCALE serves as an effective in silico tool for large-scale protein-RNA binding {Delta}{Delta}G prediction, which pushes the border of understanding in mutation-induced pathological outcomes.

20
Mapping genome-wide RNA-RNA and RNA-DNA interactions in nuclear hubs of male and female Drosophila cells

Gunasekera, S.; Carlson, M.; Ray, M.; Larschan, E.

2026-08-06 genomics 10.64898/2026.08.02.742296 medRxiv
Top 0.8%
2.4%
Show abstract

Nuclear bodies are nucleoprotein complexes with established functions that target chromatin at specific locations and regulate specific RNA processing functions, thereby influencing gene expression. However, the mechanisms that define how nuclear bodies are targeted to specific locations within the genome where they function remain poorly understood. One significant challenge is capturing and understanding the multiple cell-specific interactions occurring in these complexes, arising from RNA components interacting with each other and with DNA and nucleic acid-binding proteins within the context of the nucleuss three-dimensional organization. Mapping these interactions is critical for elucidating mechanisms such as RNA splicing, a key driver of cell-specific transcript diversity. Here, we use RNA-DNA Split Pool Recognition of Interactions by Tag Extension (RD-SPRITE) to characterize, for the first time, sex-specific RNA-RNA and RNA-DNA interactions in Drosophila S2 (male) and Kc (female) cells. We determined the sex-specific RNA-RNA interaction map within the nucleus and, using RNA-DNA interaction data, pinpointed the target loci of various RNA molecules, including small nuclear RNAs (snRNAs), which are core components of the spliceosome-a ribonucleoprotein complex involved in RNA splicing. Based on RNA-RNA interaction data, we also identified novel long non-coding RNAs that may regulate splicing. Furthermore, we investigated the role of transcription factor (TF) CLAMP in sex-specific targeting of the spliceosome. We generated RD-SPRITE datasets in the presence and absence of CLAMP, a key TF involved in dosage compensation, sex-specific RNA splicing, and chromatin organization. We determined that CLAMP regulates global changes in spliceosomal interactions with chromatin, inhibits aberrant snRNA interactions, and regulates sex-specific interactions of RNAs involved in splicing function. Additionally, our dataset provides a valuable resource for investigating additional processes, such as miRNA-mediated silencing, nucleolar functions of snoRNAs, and Cajal body functions of scaRNAs, among others. To facilitate broad community use, we have developed a computational platform, "FlySprite," that enables Drosophila researchers to explore sex-specific RNA-RNA interactions, as well as DNA targets of RNA clusters, through a user-friendly interface.